Tech Note | VisionSuite Commissioning Guide

Master the commissioning process for VisionSuite to optimize performance and ensure seamless integration in your tech projects.

Updated at July 31st, 2026

Information


This guide will walk through requirements and steps to ensure a proper functioning VisionSuite system.

Requirements

Software

  • Q-SYS Designer: Version 10.4.1 or later
  • VisionSuite Accelerator: Version 119.1.0 or later (public patch download will be here)
  • MXA920 firmware: Version 6.8 or later
  • TCC2 firmware: Version 1.8.10 or later
 
 

Hardware

  • Q-SYS Core 24f or better
    • Core Nano, Core 8 Flex, NV- 32 (Core Mode) and Core 110f are not supported
  • Q-SYS VSA-100
  • Presenter Spotlight / Speaker Camera:
    • NC-12x80
    • NC-20x60
  • Conductor Camera:
    • NC-12x80
    • NC-20x60
    • NC-90-G2
    • NC-110
  • Supported Micropohone:
    • Shure MXA920 (recommended) up to 8
    • Sennheiser TCC2 up to 8
 
 

Tools & Materials

 
 

Site Readiness

To ensure a smooth VisionSuite commissioning, the rest of the space should be fully functioning by the time VisionSuite is installed. Everything (even furniture!) in the room should be mounted, set up, and ready to go. The camera automation configuration is the last step in an integrated AV space commissioning workflow:

  1. NC Series cameras, microphones, Q-SYS Core, VSA-100, etc should already be installed, connected, and configured
  2. The network should be properly deployed
  3. Complete DSP and control commissioning should be completed (excl. VisionSuite)
  4. The appropriate licenses should have been loaded and installed
  5. The room should be free for an entire business day (not in production or business use)

Update VSA

Use the steps below to ensure VSA hardware is running the correct firmware:

  1. Open Q-SYS Designer
  2. Add the desired Core processor Status component to the design
  3. Save to Core and Run to update the Core firmware
  4. Disconnect the design
  5. Add the VSA-100 component as a placeholder
  6. Configure the IP for your discovered VSA
  7. Save to Core and Run
    1. This will update the VSA to the matching firmware, with the installation progress visible on the VSA component in real-time
  8. Download the latest update from the Q-SYS VisionSuite Updates Page
  9. On the Core, navigate to Patch Manager (http://CORE_IP/patch/vsa_patch)
  10. Upload the downloaded firmware using the Local upload button
  11. Select the patch
  12. Click the Install button
  13. Give the core and VSA 15-20 minutes to handle the installation (speeds vary depending on the network)
  14. Progress will be visible in real time in Patch Manager
 
 

Phase 1: System Design

Understand System Application

Note

Native privacy and auto privacy in Q-SYS USB Bridging devices is not currently supported.

 

New Design vs. Upgrading from VS Legacy

  • In most cases, it is possible to completely reuse audio, video and non VisionSuite control processing when upgrading your design. Things to consider are:
    • Control programming that interfaces via named components with ACPR and Seervision plugins
    • Custom scripts that interface directly with Seervision’s public websocket API
    • Advanced programming that enabled room division use cases
    • Control scripts for UCIs to enable automatic and manual modes
    • In general, most custom programming will have to be reimplemented when upgrading an existing VS Legacy system - this will take additional time. Please see the best practices guide for highlevel time estimates and how to update an existing system here: <need link to new page>

UC Compatibility

  • Using the Q-SYS Connect component there are several UC platforms that are drop in compatible with VisionSuite. Standard UCI and control programming can be used to incorporate VisionSuite seamlessly with your UC platform of choice:
    • Microsoft Teams Room
    • Zoom Rooms
  • Incorporate Call Sync into your VisionSuite system to seamlessly show status of the system on NL-SB42 speakers and TSC Generation 3 touchpanels
  • Additional functionality may require advanced Lua scripting to extend the base feature set of VisionSuite in your room.
  • Incorporate as many Q-SYS USB bridging and Mediacast to HDMI endpoints as necessary. Only a single Mediacast output from the VS component is supported at this time, but you can implement any “one-to-many” applications as needed.

Plugins

  • Seervision and ACPR plugins should not be used alongside the VisionSuite component. VS Native is not compatible in the same design with VS Legacy applications.
  • Several 3rd party plugins can be used with VisionSuite systems as is standard in Q-SYS installs. Be cautious of adding several process intensive plugins to any core serving as the connection point to VisionSuite Designer. Ensure you maintain a safe compute and control usage during normal runtime. Advanced statistics are observable by changing the core status component property Verbose = Yes.
 
 

Create Design File

  1. Open Q-SYS Designer and create a design file:
  2. Add all hardware to your design:
    1. Add a VSA-100
    2. Add your cameras
      1. Camera properties
        1. Set IP streaming to Multicast Only
      2. Camera controls settings (in QSYS design at runtime under “Status”)
        1. Set Flow Control to Off  
        2. Set Traffic Shaping to On
    3. Add a VisionSuite component
    4. Wire your cameras to the Mediacast inputs of the VisionSuite component

      Note

      Cameras must be wired in the schematic and must have the correct Input number to use the Ceiling Mounted checkbox. Each camera must have a unique Input number. This also applies for emulation.

       
    5. Wire the VisionSuite component’s Mediacast Output 1 to one more Mediacast destinations in any “one-to-many” configuration. Mix and match multiple USB video bridges or NV-based Mediacast to HDMI as needed for your application.
  3. Add all microphone channels into your design:
    1. Optional: Add the MXA920 / TCC2 plugins into the design (Q-SYS Library)

      Warning

      Sending AEC reference back to the TCC2 will result in a talker position directly below the microphone when the far end is talking. In most cases, this results in a poor user experience and is not recommended.

       
    2. Configure audio routing and processing for the microphones.

      Caution

      Avoid on-board microphone processing as much as possible. Certain on-board microphone processing will delay the audio stream but not the talker’s coordinate data. This can cause mismatches in audio received vs. audio detections.

       
  4. Design your audio processing and understand your room acoustics
    1. Understand any acoustic challenges or sonic issues in your space before you design the audio processing. If possible, apply physical acoustic treatment to the room to reduce audio reflections and acoustic anomalies to improve the effectiveness of all stages of audio processing.
    2. VAD Input
      1. Design your VAD processing to be responsive and low noise. The VAD input must be broadly intelligible and well optimized for speaker detection. We recommend utilizing a Gating Automatic Mic Mixer for best performance.
      2. With no talker audio, verify the meters show ~0 dB signal-above-noise when nobody speaks. If they don't, check upstream gain staging. For scenarios with constant noise / HVAC, apply more noise reduction in the in the AEC component as needed.
      3. Have someone speak at normal volume; the gate should open cleanly. Whisper; decide if you want whispers to trigger the gate. Adjust threshold accordingly.
      4. Have someone speak with natural pauses. If the gate closes while the person pauses, increase hold. If it feels sluggish switching to another speaker, decrease it.
      5. Depth = 60 dB.
      6. Last Mic On = Enabled
      7. Max NOM = 1
        1. Typically ensures only one mic is correlated with the incoming coordinate data
    3. Speakerphone Out
      1. Design your Speakephone Out processing to be natural and transparent. We recommend utilizing a Gain-Sharing Automatic Mix Mixer.
      2. With no talker audio, note the ambient level on meters, then set threshold ~6 dB above that.
      3. Start at Depth = −15 dB. Have one person speak; listen for background noise from inactive channels. Increase depth only as needed.
      4. Have two people alternate speaking. If the handoff sounds choppy, increase hold time. If it feels sluggish, decrease it.
      5. Leave Attack at 5 ms unless you hear plosive artefacts.
    4. Speakerphone In
      1. Ensure the audio level is appropriate and the gain in the speakers is loud enough to be clearly heard in the room. Minor EQ adjustments may be applied before the speaker output if the room has acoustic challenges.
    5. Far-End Input
      1. Design your Far-End processing to be responsive and low noise. The Far-End input must be broadly intelligible and well optimized for speaker detection. We recommend utilizing an additional gain component to ensure proper audio level enters the VisionSuite component.
    6. Due to issues with excessive frame freeze when cameras are switched, NV-21-HU USB bridges are not recommended at this time.
    7. Only one Mediacast output is supported per VisionSuite component at this time.

For more information, see Q-SYS Quantum Level 1 Training (Online).

 
 

Configure Room in VisionSuite

In emulation mode, double-click the VSD component to open the VisionSuite Designer interface.

  1. Create a Room and create a Config, giving them friendly names

  1. On your floor, upload the floorplan and click ‘calibrate scale’
    1. Place your blue dots at a known measurement and type in the actual distance to calibrate

      Tip

      Once you’ve imported the floorplan, move it by clicking and dragging so that the origin (0,0) of the project starts at the bottom left corner of the room’s floorplan. This way, when you start placing devices and furniture, their values will all be positive relative to the origin.

       

  1. On the room, set the room dimensions via the room zone and under Room shape set the room height (min and max).
  2. Go to Inventory and add all your devices:
    1. Add all cameras, and if known, their rough location in the room.
      1. Specify each camera’s type:
        1. Speaker Spotlight (Audio Driven)
        2. Presenter Tracking (Vision Driven)
        3. Conductor (Vision Driven Overview)

          Note

          Up to 14 Audio Driven (Speaker Spotlight) cameras and 1 Presenter Tracking, and 1 audience-facing conductor camera is supported.

           
    2. Add all ceiling microphones

      Note

      Up to 8 ceiling microphones are supported.

       
    3. Add the VisionSuite Accelerator
    4. Configure each device's properties accordingly

  1. Optional: Add furniture and displays into your room to have a 3D reference of the layout.
 
 

Configure Shot and Position Containers

Overview

There are two primary container types, depending on the automation mode: Shot Containers and Position Containers (Presets).

Position Containers define where the camera goes, using fixed pan, tilt, and zoom (PTZ) values to always produce the same static view, typically used for overview or preset shots. Shot Containers, on the other hand, define how a person is framed, dynamically adjusting the camera based on the subject’s position to maintain consistent composition (e.g., centering and sizing the speaker).

To access them:

  1. Select a camera
  2. Locate the right-hand panel
  3. Under the video preview, open the Containers tab

Shot Containers

Note

Shot containers that do not include a position component are shared across all cameras

 

You can create shot containers in two ways:

  • Click the “+ Shot” icon to create a default shot container
  • Click the dropdown arrow next to “Shot” to select a predefined shot with a specific relative size
  • After taking a “Shot” you should see the following container:

Shot Container Components

Each shot container includes several configurable parameters, which can be edited by right-clicking the container and select ✎ Edit. The most important are outlined below.

Transition

  • Settle Time - Defines how long it takes for the shot to complete after a transition begins. There is trade-off between a short and a longer settle time. Having settle time less than 1 sec may provide with a faster switching at the cost of accurate framing. Giving more time to the system (i.e. > 1sec) to settle, improves accuracy of framing.

Shot

  • Framing X - Controls the subject’s horizontal position within the frame.
  • Framing Y - Controls the subject’s vertical position within the frame.
  • Relative Size - Defines the subject’s height relative to the image height. In practice it defines how zoomed-in the subject should appear in the frame.
    • A value of 0.4 (minimum) results in a wide shot, where the subject appears small relative to the image.

  • A value of 3 (maximum) results in a tight shot, where the subject appears large in the frame (highly zoomed-in).

Shot Safety 

Defines how much margin (buffer space) the system keeps around the subject to ensure they remain properly framed. This margin is derived from the covariance of the audio detections (does not apply to continuous vision-based tracking).

Note

This setting is primarily relevant when using microphone-only data, where subject positioning is approximate.

 

 

In practice, it compensates for inaccuracies in microphone-based positioning, preventing the subject from being partially out of frame.

  • Lower values → tighter framing, higher risk of cut-off
  • Higher values → safer framing, more zoomed-out shots

Relation to covariance (detection uncertainty):

  • Low covariance → accurate position → tighter framing is safe → lower Shot Safety
  • High covariance → uncertain position → wider framing needed → higher Shot Safety

Speaker Height

For any zone using a container with a fixed Speaker Height, all subjects within that zone will be framed at the specified height.

Note

Does not apply to continuous tracking, and should be used mainly for TCC2 mics.

 

In practice, this locks the subject’s vertical framing, preventing people from being partially cut off.

Tracking

  • Continuous (Tracking Shot) - this option should be used primarily for Presenter Spotlight use cases. The system continuously adjusts the shot to follow the subject’s movement.

  • Frame (Static Shot) - this option should be used primarily for Speaker Spotlight use cases. The system moves to a target framing and stops. The shot remains static once the target position is reached.

 
 

Position Containers (Presets)

Note

Position containers are camera-specific and cannot be executed by other cameras.

 

Position containers define explicit pan, tilt, and zoom (PTZ) values for a camera. Unlike shot containers, they represent fixed camera positions rather than dynamic framing rules.

They can be easily identified by their thumbnail preview, which reflects the camera’s view at the time the position was created.

You can create position containers by clicking the “+ Position” icon. The container will capture the camera’s current pan, tilt, and zoom values.

Position Container Components

The primary configurable parameters are:

  • PTZ (Pan, Tilt, Zoom)
    • Pan: Horizontal rotation of the camera
    • Tilt: Vertical rotation of the camera
    • Zoom: Magnification level of the camera. These values determine the exact viewpoint of the camera.
  • Transition
    • Settle Time - Defines how long it takes for the camera to move from its current position to the target PTZ position.

Note

Position containers can be created in emulation mode, but will need to be updated once on-site to save the camera position once connected to the system.

 
 
 
 
 

Depending on the automation, different container types and container settings are optimal

Container Settings by Use Case

To customize or make changes to your containers, right-click the container and select ✎ edit. You might have to click ‘+ Add component’ at the bottom of a container to add some of these settings.

Single Presenter Tracking

For single presenter tracking, you’ll want a shot container with the following settings:

  • Settle Time: ~3s
  • Relative size: 1.3 - 1.8 (Half-body)
  • Tracking: Continuous
  • Deadband: Enable
  • Tracking Smoothness: 0.5

Tip

To add these settings, simply click ‘add Component’ and search for the variable.

 

Note

  • Remove the relative size component and set a fixed zoom instead to mitigate constant reframing.
  • Disable the Tilt axis component if the presenter will only move left-to-right.
 

Single Speaker

For Speaker Spotlight, you’ll want a shot container with the following settings:

  • Tracking: set to Frame
  • Transition Settle Time: 0.5s

    Note

    The camera will never move this fast. Trajectory velocities will be adjusted to be as fast as possible.

     
  • Shot safety: 0.3 (default, increase later to 0.5 - 1 if you see subjects being cut off)
  • Speaker height: Fix to ~1.2m (if participants will always be seated – very recommended!)
  • Relative size: 1.3 - 1.5

Note

  • Speaker height should be used with TCC2 microphones. When using Shure microphones it is not needed.
  • Remove the relative size component and set a fixed zoom if participants will all be at the same depth from the camera. Note that these containers are not ‘global’ and will have to be created per camera.
  • If participants will always be seated on the same plane, remove the Y-axis framing to eliminate height variations in the shots.
 

Tip

  • Shot safety is suitable at close distances (at most 4-5m / 13-16ft from mic). If a large relative size is used, we recommend pairing this with an equally high shot safety. This will result in more zoomed-out shots on average, but you will most likely get decent framing.
  • You can drag the framing skeleton or framing disc in the preview to adjust the headroom.
 

Home/Silence Position

For triggering a preset container whenever there’s silence or the far-end is speaking, you’ll want to create a position container. Ideally, you’ll want to have a dedicated camera for the overview, but this can also be created on one of your Speaker Spotlight cameras as well.

This container should have an overview of all the participants, providing an ‘establishing shot’. 

 
 

Note

  • Since the Z-axis data is less accurate, not having a fixed speaker height in your containers will mean more vertical variation in shots, especially when amplified by acoustic reflections from tables or windows.

  • The bigger the shot safety value, the wider a shot will be when taking into account the position uncertainty of a talker’s location. At the same time, the further away a talker is from a microphone, the higher the location uncertainty will be.
 
 
 

Create Speaker Zones

Overview

Zones define:

  • Where VisionSuite focuses attention
  • What areas are ignored
  • How cameras respond to detected activity

Tip

Speaker Zones can be initially configured in Emulation Mode and later refined on-site during commissioning.

 

Speaker Zones

Speaker Zones monitor audio activity within a defined area. When speech is detected inside a zone, VisionSuite can trigger a corresponding camera action.

Best Practices

  • Place zones at speaking locations
    • Cover areas such as seats, presentation spots, or audience sections.
  • Use generously sized zones
    • Avoid small or tightly defined zones. Tight boundaries can lead to unstable talker detection, causing delayed or erratic camera switching.

Zone edges should not align exactly with seating positions. Instead, include a buffer around each subject area to account for natural movement (e.g., people shifting in their seats) and microphone localization variance.

This helps ensure that:

  • Speakers remain within the zone even as they move
  • Minor detection drift does not trigger incorrect zones or missed activations
  • Account for coordinate drift
    • Microphone positioning is approximate (±20–30 cm), not exact. Slightly larger zones help maintain reliable detection.
  • Plan camera coverage
    • Assign zones to cameras that have clear sightlines, ensuring natural, frontal shots when triggered.

Creating Speaker Zones

  1. Open the Speaker Zone on the left sidebar
  2. Click the “+” (Add) icon
  3. Select the newly added zone and on the right sidebar the configuration of the zone should be visible:
  4. Select the Trigger type on the right sidebar
  5. Position the zone by adjusting its corner points or by enabling the Move entire zone option and moving the entire zone shape
  6. Assign a descriptive name

Tip

For custom shapes beyond trapezoids, double-click any corner to add more vertices to the zone shape to accommodate irregular layouts. To delete a corner, select it and press delete on keyboard.

 

Note

You’ll want a Speaker Zone even for the presentation area. The Speaker Zones will be responsible for switching the camera feed when the presenter speaks, and then your Presenter Zones can kick in to automate the presenter tracking.

 

Example Zone Layouts

U-shape Meeting Room

All-Hands Event Space

Oval Boardroom

Lecture Hall

 
 

In addition to standard Speaker Zones, there is another type available under Type:

Spotlight Exclusion

The Spotlight Exclusion zone is used to ignore audio detections within a defined area. Any speech or noise originating inside this zone will not trigger speaker-based rules or camera actions.

This is particularly useful for filtering out:

  • Persistent background noise sources
  • Audio reflections or reverberation artifacts
  • Areas where localization is unreliable

Example

In a room with an oval table, detections may drift toward the center due to microphone approximation. Placing a Spotlight Exclusion zone in the center prevents false triggers from that region.

Presenter Zones

Presenter Zones are used for camera behaviors related to presenter tracking and framing. Unlike Speaker Zones, which rely on audio activity, Presenter Zones are primarily driven by visual tracking and camera behavior.

  • They can be created in Emulation Mode
  • Their shape can only be finalized on-site, once the system is connected to the cameras

Note

Presenter Zones are only applicable to Presenter Spotlight or Conductor cameras.

 

Presenter Zone Types

There are three types of Presenter Zones:

  • Trigger Zones
  • Exclusion Zones
  • Tracking Zone

Detailed configuration and behavior for these zones are covered later during the on-site setup phase.

 
 

Create Speaker Rules

Overview

Rules define how camera automation behaves in VisionSuite.

A rule links:

  • An Event (e.g., activity within a zone)
  • To an Action (e.g., triggering a container, switching cameras)

In essence:

Event → Action

Prerequisites

  • Ensure Speaker Zones are defined
  • Ensure cameras are logically assigned to cover each zone

Rules can be created from the right-hand panel after selecting your configuration.

Example Setup

Speaker Logic

To trigger a camera on an active speaker, we’ll first focus on the ‘Speaker Logic’ section, creating individual rules for all of our Speaker Zones.

Rule: One Person Speaks in One Zone

This is the foundational rule for speaker-based automation.

Configuration Steps

  1. Click the “+” (Add Rule) icon on the right of the Speaker Logic section

  1. Select a Secondary Camera
    1. Example: Left Secondary PTZ
  2. Assign the same Shot Container

This allows the system to:

  • Alternate perspectives between adjacent speakers
  • Minimize noticeable camera repositioning

Without a secondary camera:

  • Transitions between nearby speakers may appear abrupt
  • PTZ movement becomes visible and distracting

With a secondary camera:

  • Switching feels smoother
  • The system can “cut” instead of “move”

Tip

You can add as many cameras as needed to a particular rule. It is not limited to primary and secondary, simply click ‘Add Camera’ to do so.

 

Once the first ‘single zone speaker spotlight’ rule has been created, repeat the same for the rest of your Speaker Zones. In this example, we’ll repeat the same for the left side of the U-shape and for the presenter area. Make sure to enable the zone by turning on the toggle switch.

Rule: Nobody Speaks in Any Zone (Silence Rule)

This rule defines the system’s fallback behavior when no active speaker is detected.

Configuration Steps

  1. Create a new rule
  2. Set Event Type to “Nobody Speaks in Any Zone”
  3. Enable “Switch Camera”
  4. Set Hold Time to approximately 5 seconds
    1. This ensures the system waits for sustained silence before triggering
  5. Select the Overview Camera or a Speaker Spotlight camera with a “overview” Position Container

Behavior

After a period of silence:

  • The system switches to an overview shot
  • This provides a neutral, stable view of the room

Note

  • If no dedicated overview camera is available, you can use a home position on a Speaker Spotlight camera.
  • Avoid very short hold times here, as they can cause rapid switching during natural pauses in conversation.
 

Rule: Multiple People Speak in One Zone

This rule handles overlapping speech within a single zone.

Use Case

When multiple participants in the same area speak simultaneously, a single-speaker shot may no longer be appropriate.

Recommended Actions

  • Trigger a group (static) shot that frames all speakers
  • Or switch to an overview camera

Rule: Multiple People Speak in Multiple Zones

This rule applies when simultaneous speech occurs across different zones.

Use Case

Cross-room discussion (e.g., interruptions, back-and-forth dialogue).

Recommended Actions

  • Trigger a group (static) shot that frames all speakers
  • Or switch to an overview camera

Practical Insight

Handling multi-speaker scenarios correctly:

  • Prevents erratic camera switching
  • Maintains viewer context during dynamic conversations

Rule: Far-End Speaking Spotlight

This rule manages behavior when remote participants (far-end) are speaking.

Objective

Ensure in-room viewers (and recordings) provide contextual visibility instead of staying locked on the last in-room speaker.

Configuration Steps

  1. Create a new rule
  2. Set Event Type to “Far-End Speaks”
  3. Set Hold Time to approximately 5 seconds
  4. Enable “Switch Camera”
  5. Select the Overview Camera or a Speaker Spotlight camera with a “overview” Position Container

Behavior

When far-end participants speak:

  • The system transitions to an overview shot of the room
  • This provides spatial context and avoids misleading framing
 
 
 
 

Phase 2: Connecting & Calibrating

Configure Microphone Arrays

Once the microphones are installed in the room, it is important to configure them properly in their native interfaces.

Shure MXA920

In Shure Designer, see Shure Best Practices.

Set Device Height

  1. Navigate to Coverage → Properties → Position
  2. Enter the precise Z-height (distance from floor to microphone center) in meters

Note

  • An incorrect Z-height will cause the microphone to report wrong speaker positions, which cascades into bad camera framing
  • Best performance is achieved with ceiling heights below 2.5 m, although Shure’s maximum height recommendation is 3.7 m
 

Enable Automatic Coverage

  1. Keep Automatic Coverage enabled, it dynamically manages lobes for better speaker tracking.
  2. Keep Auto-Focus enabled
  3. Disable on-board processing on the microphone itself to prevent conflicts in the intelligibility of talkers

Note

  • Only use manual lobe configuration for special cases (e.g., fixed podium positions)
  • If using manual lobes, remember to set Z-height for each lobe individually
 

Configure Coverage Areas to match Room Layout

  1. Place coverage areas where people actually sit, not over areas where no activity is expected
    1. For best performance, keep coverage areas small (around 5x5m)
    2. For multiple microphones keep coverage area overlap minimal (<20%), and for each part of the room, ensure the closest microphone is covering that area
  2. Place exclusion (muted coverage) areas over noisy parts of the room such as doors, windows, or HVAC

Note

If microphones are reporting detections outside the configured coverage areas a factory reset can resolve this issue.

 

Configuring Reporter and Follower Microphones

Note

Turning this feature on will cause the Shure mics to fix the height of all detections. Therefore it only works for rooms with seated participants.

 

With the release of Shure’s firmware v6.6 for the MXA920, the ‘Optimize Tracking’ feature enables aggregated sending of talker position data.

This feature enables one ‘reporter microphone’ to send talker position data from up to three additional ‘follower microphones’, supporting a total of 4 microphones in each ‘camera tracking group’.

To configure this:

  1. Install all microphones so that the status LED face in the same position for each mic
    1. Don't install microphones more than 3.6m (12 ft) from each other
  2. Place the microphones in the Shure Designer room to match the positions of how they're installed
  3. Enter the precise height for each microphone in Shure Designer (Coverage > Properties > Position)
  4. Open up the “Online Room” in Shure designer
  5. Go to the room’s Coverage view
  6. Click “Optimize tracking”
  7. Add only the reporter microphone in Vision Suite Designer

Test Microphone Pickup

  1. Have someone speak from different positions in the coverage area
  2. Verify the microphone is receiving audio

Configure Ceiling Microphone IP in VisionSuite Designer

  1. Enable the ‘Microphones Raw Data' visualization in the Visualizations dropdown
  2. Verify that you can visualize talker locations, even if they’re not yet accurate due to lack of calibration

Configure Mics in VSD

  • In a highly reverberant room, turn on the “Reflective Acoustics” global parameter in Vision Suite Designer, to optimize the mics reflection/height correction for the space
  • In cases of high background noise (HVAC, server noise etc), turn on the “Background Noise” global parameter in Vision Suite Designer, to optimize the mics vad sensitivity and localization sensitivity settings for the space
 
 

Sennheiser TCC2

Configure TCC2 Exclusion Zones

  • Use exclusion zones to block doors, HVAC, projectors, and other predictable noise sources from sending audio/coordinate locations
  • For rooms with multiple TCC2 units:
    • Only use exclusion zones for potential unwanted noise sources, not to create dedicated coverage areas in a multi-mic case. If a speaker is in a TCC2 exclusion zone and are loud enough the mic will unfortunately report a detection from the edge of the zone.
    • VisionSuite Designer will automatically merge or exclude tracking information from multiple TCC2 units based on its own internal logic

Validate Audio Pickup before linking to VSD

  • Test each TCC2 individually:
  • Have someone move and speak around the room
  • Confirm that Control Cockpit shows audio pickup and a heatmap or position indicator
  • Proceed once the TCC2 behaves correctly

Add the TCC2 devices to VisionSuite Designer

  1. Add each microphone’s IP address in the VisionSuite Designer inventory
  2. Enable Visualizations → Microphones Raw Data
  3. Confirm Talker location indicators appear in the visualization
  4. Confirm all TCC2 units are communicating with VSD
  5. Verify that you can visualize talker locations, even if they’re not yet accurate due to lack of calibration

Configure Mics in VSD

  • In a highly reverberant room, turn on the “Reflective Acoustics” global parameter in Vision Suite Designer to optimize the mic’s treatment of detections in this environment
  • In cases of high background noise (HVAC, server noise etc), turn on the “Background Noise” global parameter in Vision Suite Designer, to optimize the mics sensitivity threshold setting for the space
 
 
 
 

Room Calibration

Room calibration is the foundation of VisionSuite Features performance. This process helps cameras and microphones understand where they are in relation to each other and to the room, so the system can combine their data and accurately locate people. Fortunately, this doesn’t require manual measurements. Calibration uses ArUco markers - QR-code-like patterns - that let cameras determine their position just by viewing the marker. With a little user assistance, the cameras can also calibrate the microphones by observing them (more on that later).

Calibration Process

  1. Place one printed out Aruco Marker near your room origin.
  2. Point a camera at that marker and let it determine it’s location based on that marker.
  3. Use the same camera to auto-calibrate your microphone.
    1. If the camera can’t see the mic properly, or just at a flat angle, measure the location of the mic manually and fill in this info.
    2. In any case, measure the mics height and input that value manually
  4. Validate the detection-matching of this camera-mic pair.
  5. Place a marker central in the “working area” of the cameras that cover this mic’s working area. In most cases it should be placed flat at the center of your table:
  6. Calibrate all cameras with the central marker.
  7. One by one, validate the detection-matching of each camera with the mic data.
  8. Repeat the steps 3-7 for all additional Microphones and Cameras that you might have. As you add the mics, isolate them by disconnecting all other mics in VSD and test the performance to make sure THAT mic works well. (change port no. or IP)

Do

  • Measure the position of the first marker precisely. The placement of this marker determines how well the placement of the cameras and mics matches the real world, which will make troubleshooting so much easier.
  • Enter the correct marker size when you add the markers to VSD.

  • Place your markers so that they are clearly visible from the cameras. They should be at a 30-60° angle towards the cameras.
  • Use an angled surface (TSCs are great for this) to get a better viewing angle for your cameras, if your calibration seems to fail:

  • During calibration use a medium zoom level that corresponds to the average zoom during operation.
  • Use the grid in the 2D/3D-View to sanity-check the positions of your inventory items. Each light gray grid tile is one meter.
  • Use manual measurements to validate the positions of your inventory items.
  • Enable person data on the camera streams to compare mic data and camera data.
  • Add more markers to the room to validate the cameras.
  • Speak towards the mic and avoid speaking close to reflective surfaces when testing. You want to test the quality of your calibration, not the system performance.
 

Don't

  • Use manual measurements to calibrate your cameras, markers and mics.
    • Exception: The first marker and the mics IF they are at an awkward angle.
  • Place your markers full frontal or at very low angles towards the camera’s. Instead stay between 30-60°.
  • Place the secondary marker’s outside of the “working area” of your cameras - you want to optimize for where the cameras will be looking during use.
  • Place your secondary marker(s) in front of the room, close to your cameras, or in the far back. The center is the sweet spot.
  • Use multiple markers for the same working area. You want a single marker as the same source of truth for all cameras that cover this area.
  • Place a marker on an object that might move during calibration.
 

Calibrating Camera with Marker

  1. Enter “calibration mode 📐”  under your Room tab
  2. Select a camera
  3. In the ‘Select microphone or marker' dropdown, choose the marker that you want to calibrate or use to calibrate the camera
  4. In the camera preview window, point it at the marker the marker is centered in the frame
  5. Click ‘Calibrate'
  6. Click ‘Change the placement of’ → Marker or Camera, depending on which one you want to change
  7. If the position looks realistic, click “apply”, otherwise “view alternative result”
  8. If the alternative result looks realistic, click “apply”
 
 

Calibrating Microphone - Vision-based

  1. Enter “calibration mode 📐”  under your Room tab
  2. Select a camera
  3. In the ‘Select microphone or marker' dropdown, choose the microphone that you want to calibrate
  4. In the camera preview window, point it at the ceiling microphone so that it is centered in the frame
  5. Click ‘Calibrate’, and place the yellow box overlay to align with the shape of the microphone
  6. Click ‘Confirm corners’
  7. Drag the white square on top of the LED light of the microphone
    1. For Sennheiser TCC2 microphones the LED you want to select is the one to the left of the Sennheiser logo, BUT this will only match if the microphone is in its factory default orientation in the TCC2 Cockpit software (LED in the top left)  

  1. Click ‘Confirm LED location’
  2. Click ‘Change the placement of’ → Microphone
    1. If the microphone’s position and height look realistic, click ‘Apply New’
 
 

Validating Mic & Marker Position

  1. Enter “calibration mode 📐”  under your Room tab
  2. Select a camera
  3. In the ‘Select microphone or marker' dropdown, choose the microphone or marker that you want to validate
  4. Click “validate”
  5. The camera moves to where it thinks the microphone or marker should be. Sometimes the move doesn’t complete, click “validate” again to make sure.
  6. The microphone or marker should be roughly centered in the camera view.
 
 

Validate Detection-Matching of Mics & Cameras

  1. Enable the person data in the camera stream settings
  2. Point the camera you want to validate to look at you when you are standing in the room
  3. Speak. You should see person data popping up with a blue dot in the center and a grey area around it.
  4. Repeat this in multiple positions in the room.

The dot shows where the system thinks that your mouth is in the camera view. It may be slightly off, but if it is consistently off, the calibration is probably off. 

 
 

Best Practices & Troubleshooting

Placing Markers

  • You need to place them in the most representative area of the camera(s) that are going to be working on that area. 

  • For example if a room setup has split the audience space into 4 rectangles of 5x5meters and each audience sub-area is served by one primary/secondary camera pair, then a marker needs to be placed in the center of each sub-area.
    • Marker placement in the calibration frame: It occupies as much space in the frame as the speaker(s) that would normally be framed by this camera. We will give more details on this in next sections.. Also it is centered in that camera’s frame as much as possible.
    • Don’t fully zoom a camera in order to calibrate it. If you need to fully zoom, it means that you are too far away.
  • This will calibrate the camera optimally for the average shot that the the meeting participants each this sub-area are going to need. The further away (from the calibration spot) the camera needs to frame a person and the more the camera zoom deviates (from the zoom level that was actually calibrated), the more inaccuracies you should expect.
  • If a marker is close to two such areas and representative of both then its ideal to calibrate more cameras with the same marker to avoid calibration accumulated errors.
  • The marker will need to be calibrated using an already calibrated camera using the same principles we already discussed above.
    • Troubleshooting: if you continue observing some standard deviation in some axis,
  • The only limitations are:
    • The marker needs to be visible by the camera.
    • The marker needs to be ideally placed on a horizontal surface (like a table or on the floor). Avoid placing the marker in such a great distance that the camera-marker viewing angle is less than 5 degrees.

  • If you need to place the marker on a vertical surface (like a wall as we did with the reference marker) you need to keep that angle between 30 degrees and 60 degrees to get accurate calibration results.

If you face extreme problems withe 5 degrees restriction, you can always perform the calibration by placing a marker on an angled surface. Make sure that this surface stays in the same spot throughout the process.

 

 
 

Which Mic & Camera Pair should be used as initial pair?

Ideally you pic a pair in the middle (or the most central spot) of your setup. You want to minimize error accumulation by having too many calibration steps away from your first calibrated pair.

 
 

What is the room origin and how should the marker be placed to close to it?

The room origin is just the reference point that all inventory items are measures from. In other words, it is the (0,0,0) position in the room. Where you place it depends on your taste and the room plan, but it probably is either the bottom back left or the bottom front right corner of the room. You don’t input the world reference point anywhere in the system, but it is implicitly defined.

The constraint on your selection is that the z=0 coordinate needs to be on the floor, so that every objects height is measured from the floor.

When you place your first marker close to the room origin, place the marker perpendicular to the walls to make sure that you shouldn’t have to worry about the marker’s rotation in x or y, and just need to determine it’s z rotation, which should be a multiple of 90°.

When measuring the position of the marker in relation to the world origin, always measure from the center of the marker. 

 
 

How precise does calibration need to be? What errors can be expected?

This is pretty easy to answer for the mics and a little harder for the cameras. For the microphones, if you “misplace” the mic by 1cm, the detections will be misplaced by the same amount and assuming the maximum mic-speaker distance of 2.5m, every 1° of rotational error, misplaces a detection by 7cms.

The effects of camera position and rotation errors are different due to perspective and parallax effects. But to give you an idea, position errors also approximately translate 1:1, while the effect of rotation errors heavily depends on the camera-speaker distance. Very roughly 1° of rotational error misplaces a person by 7cms for each 2.5m distance between camera and speaker. This error quickly adds up for big rooms and the rotational error should be less than .25 degrees broadly. 

 
 

How well should positions in VSD match the positions of the real world equipment?

The position of your equipment should closely match the real world positions, but deviations of up to 30cms and 3° should be okay. 

 
 
 
 

Phase 3: Validate Speaker Spotlight & Fine-Tune

Validate Rules

Once the cameras and microphones are correctly calibrated, and you’ve set up your Speaker Zones and corresponding rules as covered in the ‘off-site’ section, it’s time to verify that the logic works.

In this step, speak from different positions in the room, making sure to test different seats in all of the different Speaker Zones created.

Make sure to also test multiple speakers simultaneously and people interrupting each other.

If everything has been properly set up, you should see the active speaker being framed in your desired shots. In this step, you can tweak your shot containers, including framing, shot safety, settle time, and speaker height to better meet your needs now that you see how the shots are applied.

Under the visualizations tab in the left-hand side panel, you can enable the visualizations of the detections to see the person uncertainty, the microphone raw data, and the camera raw data and get a better understanding of what VisionSuite is seeing and hearing relative to the space.

Note

Sometimes the switching speed of cameras may be slower when the VisionSuite Designer interface is opened.

 

Tip

At this stage, it's also a good time to calibrate the white balance on your cameras to ensure accurate color reproduction.

To do so, it's best to set the White Balance mode to Manual and using a 17% gray card to automatically calibrate the white balance.

For automatic white balance calibration, zoom in on the gray card using the PTZ controls (via the camera component, not on VSD), until the gray card fills up the entire frame.

Then, click the 'AWB One Push' button to initiate the auto white balance calibration. This takes around 5 seconds.

 
 
 

Global Parameters

In the Room tab (the parent of the config), you can also adjust the global settings that govern the automation in your room.

Tip

Before modifying the global parameters, double-check how accurate the room calibration is. Improving the accuracy of your calibration win ensure better performance, versus tuning these parameters.

 

Speaker Fadeout

Speaker Fadeout determines how long a person’s detection remains active after they’ve finished talking.

In the case of a longer Speaker Fadeout, if a person is interrupted, a camera may stay on the old (current) speaker for longer, waiting for the fadeout time to run out before switching the camera to the new speaker.

A Speaker Fadeout time of 2 seconds is recommended to account for short pauses when someone speaks.

If a lot of crosstalk is expected and have multi-speaker rules, set up Speaker Fadeout to above 3 seconds. At a minimum, we recommend setting it to above 1 second to avoid double switching/reframing on the same speaker.

Speaker Reactivity

Speaker Reactivity refers to the number of audio samples VisionSuite waits for before considering a speaker to be an active talker.

A lower reactivity will wait for longer to determine if a person is actually speaking. At the same time, since it waits longer, it gains higher confidence and accuracy on their coordinate positions. Nevertheless, a low reactivity will take longer to determine if someone is speaking, therefore leading to slightly slower switching speeds.

A high reactivity, in contrast, will take fewer audio samples of a speaker to determine their location and whether they’re speaking, triggering the downstream actions faster, but compromising with more inaccuracies in the confidence level and a talker’s location.

A speaker reactivity of 0.8 is recommended. If the reactivity is set to higher, the system will react before it has enough data to get a good shot. If the reactivity too low, it will reduce the switching speed too much. If that’s a compromise you’re willing to make to get significantly better shots, it is worth trying it.

Bottom line: Higher reactivity → faster switching (but could be more inaccurate)

VAD Sensitivity

Voice Activity Detection helps to mitigate false camera switches by differentiating ambient noise from speech, eliminating unnecessary distractions.

VAD operates as a speech-activated gate globally, preventing camera switching from noises only during periods of silence. Once speech is detected and the gate opens, it cannot prevent camera switching from ambient noise that occurs simultaneously while a person is speaking.

A high VAD sensitivity will be more prone to consider some forms of speech as noise, whereas a low VAD sensitivity may be more prone to consider some types of noise as speech.

You can set the VAD sensitivity value to 0 to disable VAD.

If using Shure MXA920 microphones, we recommend setting the VAD value to 0, and using Shure’s onboard VAD instead. A VAD sensitivity of 0.6 is recommended when using Sennheiser ceiling microphones. 

Far-End Peak Threshold

This parameter ensures that only audio signals higher than the set dB threshold trigger the signal of a speaker on the far-end. For example, if the far-end audio is lower than 20dB, the far-end rule would not be triggered to switch to an overview shot.

A Far-End Peak Threshold of -20dB is recommended.

Matching Sensitivity

Matching Sensitivity refers to the process of ‘fusing’ nearby detections in VisionSuite Designer.

If a microphone gets multiple detections from a single speaker, a high matching sensitivity will fuse them into a single detection (and therefore a single coordinate) to mitigate uncertainty.

However, if the sensitivity is too high, two adjacent speakers that are speaking right next to each other may be fused into a single detection and not trigger a camera to react to the new speaker.

Matching Sensitivity is therefore mostly important when using multiple ceiling microphones.

Matching Sensitivity is also applied to vision detections. If a vision camera sees a person in one location, and the audio data comes from the vicinity of that vision detection, both the visual detection and the audio detection will be merged to determine who in the camera’s view is talking.

A Matching Sensitivity of 0.4 is recommended. Above 0.9 will not be so useful, as it will most likely fuse different people into the same detection. If you see a lot of double detections, set Matching Sensitivity to 0.6. 

If using Shure’s ‘Optimize Tracking’ feature with reporter and follower microphone groups, VisionSuite’s Matching Sensitivity will be applied on top of Shure’s own clustering. Therefore, if setting low values between 0-0.4, you won't see a difference.

Reflective Acoustics

This parameter is useful if you have a room with a lot of hard surfaces and if you notice that framing often changes a lot depending on where people face when speaking. In rooms with a lot of relfections, turn the Reflective Acoustics toggle ON.

  • For Shure microphones, this optimizes reflection and height-correction handling to improve localization behavior in reflective environments
  • For Sennheiser microphones, this modifies VisionSuite’s internal modeling of detections uncertainty

High Background Noise

In situations where you have constant high ambient noise in the room, turn this toggle ON. It should not be needed in most rooms, especially if they follow our acoustic recommendations outlined in the Designing for VisionSuite guide.

  • For Shure microphones, this optimizes VAD sensitivity and localization sensitivity to improve performance in noisy environments.
  • For Sennheiser microphones, this optimizes the microphone sensitivity threshold to reduce false or unreliable detections caused by background noise.

Room Zone

The room zone defines the overall area in which the system will process detections. It should encompass all areas where you plan to create Speaker or Presenter Zones, with a small buffer around the edges to allow for movement and detection variance.

  • Typically, the room zone can match the floor outline to include the entire space.
  • Areas where you do not want automation—such as sources of noise or irrelevant regions—can be excluded from the room zone to prevent unnecessary detections.

This ensures the system only focuses on the areas that matter for camera automation.

Room Zone Height Limits

You can define minimum and maximum heights for the room zone to constrain which audio detections the system considers. Any detections below the minimum or above the maximum are ignored.

Example: Setting the minimum height to the table level prevents detections from below the table, reducing false triggers.

 
 

Validate Calibration

If performance of the system does not seem optimal, you can take these steps to validate the calibration. Make sure to turn on “raw mic data” in the settings and “persons data” in the camera streams. Make sure you are confident that you calibrated the system as instructed and do not change any settings, if not instructed to do so.

Camera Validation (QR Code Test)

Goal

Confirm cameras are positioned correctly

Steps

  1. Place Aruco Markers on both sides of the room on the wall
  2. Measure their position with a laser meter (the exact rotation is not important)
  3. Use the validation tool to check how well each individual camera frames them.
  4. Take screenshots of the framing!
  5. Repeat on the opposite side of the room

What to Expect

  • Camera on its “working side” → marker should be tracked well
  • Camera on opposite side → tracking can be less accurate, but still reasonable

If NOT OK

  • Recalibrate cameras and make sure you place the calibration marker in the center of the room.
  • Check the rotation critically. If the cameras are wall mounted it could be okay to fix the rotation at (0°,0°,0°/90°/180°,270°)

Re-test after changes

 
 

Microphone Validation (Position Check)

Goal

Confirm microphone positioning is correct

Steps

  1. Stand at a known position (measure roughly with laser meter)
  2. Speak normally
  3. Check your position in the 3D map

What to Expect

  • Position should be within ~30 cm accuracy

If NOT OK

  • Make sure that the mic is set to the correct height
  • Critically check the rotation of the mic.
  • Recalibrate the mic.

Re-test after changes

 
 
 
 

Phase 4: Configuring Presenter Spotlight

Once your Speaker Spotlight rules are working correctly, you may want to set up a camera for Presenter Spotlight. To get started, setup your containers as described in Section 1.3. Once this is done, ensure that the tracking camera in your Inventory has the following role: Presenter Tracking.

Rule Priority & Switching

Note

For a selected config, the order of rules under Presenter Logic matters. In VisionSuite Designer, rules higher up in the list override rules lower in the list. Therefore, auxiliary shots (whiteboard, multi-presenter) should be above your standard presenter tracking rule in the list. Similarly, “VIP lost” rules should be below.

 

When both modes are active, the rules engine determines which camera goes live. Rules are prioritized by type (see below)

Rule types in order of priority from highest to lowest:

  1. Far End — remote participant active
  2. Multiple Zones Active — multiple zones triggered
  3. Multiple People Talking — multiple speakers active
  4. Single Person Talking — single speaker active
  5. Multiple People Enters — multiple people entered in a zone
  6. Single Person Enters — single person entered in a zone
  7. Single Person Leaves — person left a zone
  8. VIP is Lost — tracked VIP lost
  9. No Zones Active / All Zones Silent — no zones active

In practice, this means Speaker Logic outranks Presenter Logic. If a speaker starts talking while the presenter is being tracked, the system will switch the live output to the speaker-spotlight camera.

Both containers continue to run on their respective cameras, regardless of which is live — the tracking camera keeps following the presenter even when it is not on air, so the switch back is seamless.

 
 

Trigger Zones

To enable camera automation, create a Trigger Zone by clicking the + icon next to Presenter zones in the left pane. Trigger Zones can trigger rules in Presenter Logic when activity is detected within the zone. Under Type, select Trigger. Then, select which camera the Trigger Zone should be assigned to from the drop-down menu. If required, click the toggle button to set the state to Enabled, this activates the Trigger Zone.

Use the camera video preview window to position the Trigger Zone over the area that you want to monitor for activity to start Presenter Tracking (for example, when a presenter enters). Adjust the Trigger Zone by dragging its corners.

Avoid making Trigger Zones too large, as this may cause tracking to activate when people simply pass by.

Left - Recommended | Right - Sub-optimal

Avoid placing Trigger Zones over displays or in areas such as windows or corridors, where false detections may trigger tracking. In many cases, it may not be possible to avoid placing Trigger Zones over displays. In these situations, use Exclusion Zones (see below) to prevent false triggers caused by people appearing on the displays.

Left - Recommended | Right - Sub-optimal

Tip

Presenters should naturally gravitate towards the ‘start tracking’ Trigger Zone, they shouldn’t have to know where the zone is placed to activate the tracking.

 

Once you’ve created your Trigger Zone, it is time to configure the rules that drive the presenter automation logic. In your config, head to the Presenter Logic drop-down menu and create a new presenter rule by clicking +.

All triggers are based on Enter, Leave, or VIP Lost events.

  • Events are only generated from the portion of the Trigger Zone that is visible in the camera’s view.
  • The Trigger Zone is used to take actions such as starting to track. However, leaving a Trigger Zone does not automatically stop tracking. To stop tracking, you need configure a One person leves rule or use a Tracking Zone with the VIP is lost rule. The latter is preferred, as the VIP is lost rule should always be configured to handle unexpected loss of the tracked person.

One Person Enters

The One person enters event is triggered when a presenter enters a zone and no other person is being tracked in that zone by any camera. A person is considered to have entered when a substantial portion of their head is fully within the Trigger Zone. For example, an arm or leg inside a Trigger Zone will not trigger an Enter event. For the rule to trigger, the person must remain within the zone for the duration of the hold time.

You also need to set the tracking camera and the container that is called to start tracking with the desired shot, and tick Switch Camera to ensure the camera is switched live once tracking has started.

Don’t forget to enable your rule - it won’t work otherwise. To enable it, use the toggle button next to the rule name.

Example: To start tracking, a person must enter a Trigger Zone and remain there for the duration of the hold time. Once this is met, the container (Section 1.3) associated with the Trigger Zone is activated. If the Switch Camera option is enabled, the camera feed will switch.

One Person Leaves

The One person leaves event works in a similar way and, when a camera is tracking, is only triggered when the tracked person leaves. When a person leaves a Trigger Zone, another action can be triggered, such as changing the shot size or stopping tracking. In most cases, this event is not required; instead, use the the VIP is lost event.

VIP is Lost

The VIP is lost event is not tied to a specific Trigger Zone. It is used to recall a fallback container when the presenter leaves the camera’s view - for any reason - or exits the Tracking Zone. In this case, the event is triggered when the entire body of the VIP (i.e., the tracked person) is fully outside the Tracking Zone. When this occurs, you will probably want to trigger a fallback home position container.

Multiple People are Inside

The Multiple people are inside event is triggered when the number of people visible in the Trigger zone (heads are what matters) becomes two or more. Once the number of goes back to one, the rule becomes inactive again.

Note

If you created position containers in emulation mode, make sure to update their positions with the real-life desired shot. To do this, move your camera (including zoom) to the desired position and then: right-click the container and select ‘update’.

 
 
 

Tracking Zone

A Tracking Zone helps to set the boundaries of your presentation area. It serves two functions: 1) when a presenter leaves a tracking zone a VIP is lost event is triggered that can be used to recall a fallback home position, and 2) it helps to exclude vision detections that are outside the camera's area of interest.

Tip

A person is considered to be outside the Tracking Zone when their full skeleton bounding box is entirely covered by the Tracking Zone.

 

To create a Tracking Zone, select Tracking in the zone type drop-down menu, and in the camera feed preview, drag the corners of the Tracking Zone to delineate the area where you’ll want your presenter to be tracked. Anything beyond the Tracking Zone will be considered ‘out of bounds’.

Use the Tracking Zone to also exclude any seated audience that you don’t want to interfere with the tracking behavior.

 
 

Exclusion Zones

Exclusion Zones allow you to define areas where detections are intentionally ignored by the visual logic. They are useful for eliminating unwanted detections, such as people shown on screens, and help prevent undesired camera triggers or logic events.

Exclusion Zones ignore a detection when at least 95% of the detection’s bounding box is covered by the zone. This means a tracked presenter should not be fully within an Exclusion Zone. Ensure the shot is zoomed out enough so that part of the presenter’s bounding box (for example, their legs) extends beyond the boundaries of the Exclusion Zone.

To create a Exclusion Zone, select Exclusion in the zone type drop-down menu, and in the camera feed preview, drag the corners of the Exclusion Zone to delineate the area where you’ll want your detections to be ignored.

Tip

Make Exclusion Zones generously sized and larger than the dimensions of the display. The zone may shift slightly while the camera is moving during tracking, meaning some detections may not always meet the 95% coverage threshold.

 
 
 

Audience-Facing Conductor Camera

Audience-Facing Conductor Camera Overview

The other way to incorporate vision into a VisionSuite system is as an audience-facing conductor camera.

This camera’s visual detections will aid in improving room awareness in Speaker Spotlight applications, providing contextual understanding of participant’s locations within the room, and enhancing speaker location confidence when fused with the audio coordinate data.

Such an overview camera is only practical in rooms where participants are no deeper than 7m / 23ft away if using an NC-110, or 9m / 29ft if using an NC-90.

To set up an audience-facing conductor camera, simply change your static camera’s role to ‘Overview/Conductor’ instead of ‘Tracking’. Within that camera’s video preview, we recommend placing Exclusion Zones and a Tracking Zone to ignore detections in unwanted areas and prevent fusing false detections from displays in the room.

No Trigger Zones or Presenter Logic are needed for an audience-facing vision camera.

To verify that your audience-facing conductor camera is providing the locations of participants accurately, simply enable the ‘Cameras Raw Data' in the Visualization tab to see the visual detections.

Tip

Ensure this camera’s FoV doesn’t overlap with the presenter area as to not occlude audience detections with the presenter's head (if the overview camera is not high enough). We recommend ceiling mounting the overview camera directly above the presenter area or even slightly in front of it for optimal performance. This will make sure the back of the presenter’s head does not interfere.

 

Note

Using an audience-facing conductor camera greatly improves framing consistency and accuracy, but has certain limitation especially in the number of people it can handle. The current limit is at 25 people. If that limit is reached, the system is resilient and smart enough to fall back to a home position i.e. the overview camera (can be the same as the conductor one).

 
 
 

Audience-Facing Conductor Camera Setup

The audience-facing Conductor camera provides a wide-angle overview of the room and is the default output when no speaker or zone is active. This guide covers model selection, physical placement, and VSD configuration for the Q-SYS NC-90-G2 and NC-110.

Note

Other PTZ cameras can be assigned the Conductor role but must remain completely stationary once placed. The NC-90-G2 and NC-110 are preferred because their wide fixed lenses are purpose-built for this use case. PTZ-based conductor setups are not covered here.

 

Choosing a Camera

Camera Horizontal FoV Best for
NC-110 110°

Large or wide rooms;

Ceiling-center mount

NC-90-G2 90°

Standard conference rooms;

front-center, wall or corner mount

Both are fixed-lens cameras. The NC-110's wider FoV compresses depth more than the NC-90-G2, resulting in weaker depth estimates — prefer the NC-90-G2 when depth accuracy is a priority and the room size allows it.

Wide-angle lenses also introduce radial distortion toward the edges of the frame. Subjects near the frame edges appear warped, which can significantly degrade depth estimates for those individuals. This effect is more pronounced on the NC-110. When placing the camera, ensure that the primary seating area falls toward the center of the frame rather than the edges.

 
 

Physical Placement

Presenter Position

The ideal conductor shot contains only the audience. Mount the camera directly above or slightly in front of the presenter's standing position and tilt toward the audience - the presenter then falls behind the field of view entirely. When the presenter position is unknown, mount closer to the presenter's end of the room rather than centering over the table.

Tilt Angle Trade-off

Tilt angle is the most critical placement decision:

  • Too steep (top-down): depth estimation degrades - subjects appear at nearly the same distance, collapsing the vertical parallax VisionSuite uses to separate people in 3D.
  • Too shallow (straight-on): subjects occlude each other with little vertical separation, making individual detection unreliable.

Target a downward tilt of 20–35° from horizontal at a mounting height of 2.5–3.5 m. Tilt steeper for deeper rooms, shallower for wide, shallow layouts. Keep faces visible — avoid angles where only the tops of heads are seen.

Occlusion at Max Depth

Occlusion at the far end of the room is the primary installation concern. Participants at maximum depth are most likely to overlap from the camera's perspective, and a small lateral offset near the camera becomes a large overlap at distance.

To minimize:

  • Raise the camera: height separates subjects vertically and is the most effective lever.
  • Center horizontally: distributes any remaining lateral occlusion evenly.
  • Validate at max depth: confirm heads and torsos are individually visible with people at the farthest seats before finalizing the mount.

If some occlusion is unavoidable, resolve it on the audience side where speaker detection matters most.

Mount Type

  • Ceiling front-center (recommended): best symmetry; suited to the NC-110's wider FoV. Enable the Ceiling Mounted toggle in VSD to invert pan/tilt axes automatically.
  • Wall or corner: more natural perspective; suits the NC-90-G2, but may introduce depth inaccuracies on one side of the room vs. the other.

Avoid

  • Strong backlighting (windows, bright displays) behind subjects.
  • Any position that would require the camera to move during a session.
 
 

VisionSuite Designer Configuration

Assigning the Conductor Role

Roles are set per Room Config, not globally. In the Room Config's Camera Roles panel, set this camera to Conductor. Only one Conductor is supported per config.

Validate Calibration

Navigate to the conductor camera’s inspector panel in VisionSuite Designer. Turn on the Persons data overlay. Ensure all subjects can be detected and a gray orb is drawn around their head in all typical speaking positions.

Navigate to each speaker spotlight camera’s inspector panel (primary / secondary camera). Ensure the Persons data orb is accurately projected into the speaker spotlight camera’s field of view. In each of these cameras, a subject's orb should also be drawn around their head.

If the orb is drawn too low or high across all cameras, consider recalibrating or remeasuring the conductor camera’s y rotation.

If the orb is only drawn too low or high on just one camera, consider recalibrating or remeasuring only the problem camera’s y rotation.

If the orb is too far left or right (resulting in centering inaccuracy), consider recalibrating or remeasuring only the problematic camera’s z rotation.

In general, most cameras have a near-0 x rotation.

It may be necessary to slightly adjust the calibration’s rotation results for accurate targeting.

Action Rules

The conductor camera requires no containers — it holds a fixed position and action rules switch to it directly.

Rule Trigger Action
Overview On Sjilence nobody speaks in any zone

Switch to conductor,

no container

Overview On Far-end

(recommended)

the far-end speaks

Switch to conductor,

no container

Overview On Multi-talker multiple people speak in multiple zones

Switch to conductor,

no container

The overview on silence rule is essential — it returns the output to the wide overview whenever no speaker is detected. The multi-talker rule prevents rapid cuts when several zones are active simultaneously. Avoid configuring any position containers on the conductor camera other than the fully zoomed-out default position.

 
 
 
 
 
 

Validating & Fine-Tuning Presenter Logic

Finally, it’s time to validate that your Presenter Logic works as expected.

Try whether your Trigger Zones are triggered within the defined hold time, and ensure that the Trigger, Tracking, and Exclusion Zone positions and shapes serve their functions and stay in their intended places.

We recommend running a couple of trial runs/rehearsals during actual presentations to identify any edge cases and adjust the logic afterwards to mitigate unintended behaviours.

On top of that, fine-tune your tracking shot containers, adjusting the deadband width, tracking smoothness, framing, relative size, etc to achieve your desired shots and tracking experience.

 
 

Phase 5: Maintenance & Reliability

We recommend running daily reboots of the VSA-100. This can be configured under your Room tab (parent of the config). In the ‘Schedules’ dropdown, simply select a ‘Daily’ frequency and the desired time.

We also recommend scheduling camera recalibration to take place once a week. This can also be configured in the ‘Schedules’ dropdown. Make sure to schedule recalibration during periods of time where the room is well-lit to aid in focus recalibration.

Note

Scheduling recalibration will mean that the PTZ cameras will recalibrate themselves. This is not related to the calibration of your VisionSuite Design with the calibration markers.

 

Finally, we recommend scheduling daily reboots of your ceiling microphones to maintain stable WebSocket connections with VisionSuite.

Over extended periods of time, VisionSuite may open and abandon multiple WebSocket connections to the microphone. When too many simultaneous connections accumulate, the microphone becomes unresponsive to new connection requests. A scheduled reboot clears these connections and restores normal operation.

Phase 6: Configuring Room Controls

Control Pins Configuration

To interact with the room, users will likely want to have some control over how the system operates.

In the room’s touchpanel, the following functionality toggles will most likely be needed:

  1. Switch between Room Configs (e.g. different room layouts)
  2. Bypass VisionSuite

VisionSuite Basic Controls

Both of these can be achieved by wiring Custom Controls into the control pins of the VisionSuite component.

Switch Between Room Configs

This can be used to change between different furniture layouts or room use cases, or to enable and disable aspects of the system (e.g. Presenter Spotlight only config, Speaker Spotlight only config, and Presenter Spotlight + Speaker Spotlight config).

Caution

If the system is actively tracking a presenter, switching room configs will reset all active rules and tracking will stop. It is not recommended to switch room configs during a live presentation.

 

This is achieved by sending the name of the Config you wish to recall into the Active Config control pin. The simplest way to do this is to use the Selector component.

Enter the names of the Configs, as defined in VisionSuite Designer, into the Output controls in the Selector. You can optionally add friendly label names in the label section.

You can then switch configs using the toggle buttons or the Selection combo box, and place these controls onto a UCI as needed.

Bypass VisionSuite

This will bypass both Presenter Spotlight and Speaker Spotlight.

To do this, send a boolean true/false into the Room Bypass Control Pin. The simplest method is to use a Custom Controls component with a toggle button, but you could also use a Flip-Flop or other control components depending on how you would like to interact with the bypass control.

The bypass toggle can then be placed onto a UCI.

 
 

Controlling Room with VisionSuite

Any element of a Q-SYS system can be driven by control pins. You may wish to control room lighting, audio, or any other aspect of the room, and have this automated by VisionSuite.

This can be accomplished by using the LED output Control Pins from the VisionSuite component, which are available for each room, and connecting them to other Q-SYS components, including plugins for 3rd-party devices.

In VisionSuite Designer, each LED can be programmed to become true for any event in the Config.

Automation Ideas

  1. An LED that becomes true while a Presenter Tracking rule is executing. This can be used to automatically fade out BGM and change the lighting when the presenter starts to be tracked. A second LED can be programmed to become true when the VIP is lost, which fades the BGM back in and changes the lighting again.
  2. An LED that is true every time the audience asks a question, to bring up lighting on the audience

Example

  1. In the properties of the VisionSuite component, ensure there are enough LEDs for the automation you want to perform.
  2. Create the rules in the usual way. Here there is a rule for a presenter entering the Trigger Zone, and a rule for the Presenter leaving the Tracking Zone (VIP Lost).

  1. Press the “Control Pins Configuration” button in the top bar of VisionSuite Designer.

  1. Here you will see a row for each LED.
    1. Select Rules from the Config, and a Rule Status for each LED
    2. When the status of the Rule matches the Status you specify, the LED will become true, otherwise it will be false.
Example - two LEDs, that go high at the start and end of tracking, respectively.

Rule Status Definitions

While you can select any status to set an LED, for most practical applications, it’s common to use Executing for a momentary pulse while the rule is getting from the Idle to the Live state, or Live if you want the LED to be true when the rule has completely finished.

 
 
 
 

End-User Training

When onboarding room users, make sure to explain that for the best performance, they should speak towards the direction of the microphone. Participants speaking more than 90 degrees off-axis away from the microphone will see less accurate framing. Also, verify that the switching speed and pacing of the silence rules are triggered to their liking and if not, adjust as needed.

Make sure to also explain the placement of trigger zones that are used to activate the tracking. Make them aware that moving key furniture (e.g. a lectern or whiteboard) will not move the Trigger Zone associated with that reference point. Users should also be aware of when they can expect issues (e.g. if hiding from the camera during tracking or crossing with multiple people while being tracked).